|
Corpus linguistics is the study of language as expressed in ''corpora'' (samples) of "real world" text. The text-corpus method is a digestive approach for deriving a set of abstract rules, from a text, for governing a natural language, and how that language relates to and with another language; originally derived manually, coprora now are automatically derived from the source texts. Corpus linguistics proposes that reliable language analysis is more feasible with corpora collected in the field, in their natural contexts, and with minimal experimental-interference. The field of Corpus Linguistics features divergent views about the value of corpus annotation, ranging from John McHardy Sinclair, who advocates minimal annotation, and so allow texts to speak for themselves;〔Sinclair, J. 'The automatic analysis of corpora', in Svartvik, J. (ed.) ''Directions in Corpus Linguistics (Proceedings of Nobel Symposium 82)''. Berlin: Mouton de Gruyter. 1992.〕 to the Survey of English Usage team (University College, London) who advocate annotation as allowing greater linguistic understanding, by way of rigorous recording. 〔Wallis, S. 'Annotation, Retrieval and Experimentation', in Meurman-Solin, A. & Nurmi, A.A. (ed.) Annotating Variation and Change. Helsinki: Varieng, (of Helsinki ). 2007.(e-Published )〕 == History == Some of the earliest efforts at grammatical description were based at least in part on corpora of particular religious or cultural significance. For example, Prātiśākhya literature described the sound patterns of Sanskrit as found in the Vedas, and Pāṇini's grammar of classical Sanskrit was based at least in part on analysis of that same corpus. Similarly, the early Arabic grammarians paid particular attention to the language of the Quran. In the Western European tradition, scholars prepared concordances to allow detailed study of the language of the Bible and other canonical texts. A landmark in modern corpus linguistics was the publication by Henry Kučera and W. Nelson Francis of ''Computational Analysis of Present-Day American English'' in 1967, a work based on the analysis of the Brown Corpus, a carefully compiled selection of current American English, totalling about a million words drawn from a wide variety of sources. Kučera and Francis subjected it to a variety of computational analyses, from which they compiled a rich and variegated opus, combining elements of linguistics, language teaching, psychology, statistics, and sociology. A further key publication was Randolph Quirk's 'Towards a description of English Usage' (1960)〔Quirk, R. 'Towards a description of English Usage', ''Transactions of the Philological Society''. 1960. 40–61.〕 in which he introduced The Survey of English Usage. Shortly thereafter, Boston publisher Houghton-Mifflin approached Kučera to supply a million-word, three-line citation base for its new ''American Heritage Dictionary'', the first dictionary to be compiled using corpus linguistics. The AHD took the innovative step of combining prescriptive elements (how language ''should'' be used) with descriptive information (how it actually ''is'' used). Other publishers followed suit. The British publisher Collins' COBUILD monolingual learner's dictionary, designed for users learning English as a foreign language, was compiled using the Bank of English. The Survey of English Usage Corpus was used in the development of one of the most important Corpus-based Grammars, the ''Comprehensive Grammar of English'' (Quirk ''et al.'' 1985).〔Quirk, R., Greenbaum, S., Leech, G. and Svartvik, J. ''A Comprehensive Grammar of the English Language'' London: Longman. 1985.〕 The Brown Corpus has also spawned a number of similarly structured corpora: the LOB Corpus (1960s British English), Kolhapur (Indian English), Wellington (New Zealand English), Australian Corpus of English (Australian English), the Frown Corpus (early 1990s American English), and the FLOB Corpus (1990s British English). Other corpora represent many languages, varieties and modes, and include the International Corpus of English, and the British National Corpus, a 100 million word collection of a range of spoken and written texts, created in the 1990s by a consortium of publishers, universities (Oxford and Lancaster) and the British Library. For contemporary American English, work has stalled on the American National Corpus, but the 400+ million word Corpus of Contemporary American English (1990–present) is now available through a web interface. The first computerized corpus of transcribed spoken language was constructed in 1971 by the Montreal French Project,〔Sankoff, D. & Sankoff, G. Sample survey methods and computer-assisted analysis in the study of grammatical variation. In Darnell R. (ed.) ''Canadian Languages in their Social Context'' Edmonton: Linguistic Research Incorporated. 1973. 7–64.〕 containing one million words, which inspired Shana Poplack's much larger corpus of spoken French in the Ottawa-Hull area.〔Poplack, S. The care and handling of a mega-corpus. In Fasold, R. & Schiffrin D. (eds.) ''Language Change and Variation'', Amsterdam: Benjamins. 1989. 411–451.〕 Besides these corpora of living languages, computerized corpora have also been made of collections of texts in ancient languages. An example is the Andersen-Forbes database of the Hebrew Bible, developed since the 1970s, in which every clause is parsed using graphs representing up to seven levels of syntax, and every segment tagged with seven fields of information.〔 〕 The Quranic Arabic Corpus is an annotated corpus for the Classical Arabic language of the Quran. This is a recent project with multiple layers of annotation including morphological segmentation, part-of-speech tagging, and syntactic analysis using dependency grammar.〔Dukes, K., Atwell, E. and Habash, N. 'Supervised Collaboration for Syntactic Annotation of Quranic Arabic'. ''Language Resources and Evaluation Journal''. 2011.〕 抄文引用元・出典: フリー百科事典『 ウィキペディア(Wikipedia)』 ■ウィキペディアで「Corpus linguistics」の詳細全文を読む スポンサード リンク
|